Human Mutation
○ Wiley
All preprints, ranked by how well they match Human Mutation's content profile, based on 34 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Plazzer, J. P.; Macrae, F.; Yin, X.; Thompson, B. A.; Farrington, S. M.; Currie, L.; Lagerstedt-Robinson, K.; Frederiksen, J. H.; van Overeem Hansen, T.; Graversen, L.; Frayling, I. M.; Akagi, K.; Yamamoto, G.; Al-Mulla, F.; Ferber, M. J.; Martins, A.; Genuardi, M.; Kohonen-Corish, M.; Baert-Desurmont, S.; Spurdle, A. B.; Capella, G.; Pineda, M.; Woods, M. O.; Rasmussen, L. J.; Heinen, C. D.; Scott, R. J.; Tops, C. M.; Greenblatt, M. S.; Dominguez-Valentin, M.; Ognedal, E.; Borras, E.; Leung, S. Y.; Mahmood, K.; Holinski-Feder, E.; Laner, A.
Show abstract
BackgroundIt is known that gene- and disease-specific evidence domains can potentially improve the capability of the ACMG/AMP classification criteria to categorize pathogenicity for variants. We aimed to include gene-disease-specific clinical, predictive, and functional domain specifications to the ACMG/AMP criteria with respect to MMR genes. MethodsStarting with the original criteria (InSiGHT criteria) developed by the InSiGHT Variant Interpretation Committee, we systematically addressed specifications to the ACMG/AMP criteria to enable more comprehensive pathogenicity assessment within the ClinGen VCEP framework, resulting in an MMR gene-specific ACMG/AMP criteria. ResultsA total of 19 criteria were specified, 9 were considered not applicable and there were 35 variations of strength of the evidence. A pilot set of 48 variants was tested using the new MMR gene-specific ACMG/AMP criteria. Most variants remained unaltered, as compared to the previous InSiGHT criteria; however, an additional four variants of uncertain significance were reclassified to P/LP or LB by the MMR gene-specific ACMG/AMP criteria framework. ConclusionThe MMR gene-specific ACMG/AMP criteria have proven feasible for implementation, are consistent with the original InSiGHT criteria, and enable additional combinations of evidence for variant classification. This study provides a strong foundation for implementing gene-disease-specific knowledge and experience, and could also hold immense potential in a clinical setting.
Andhika, N. S.; Biswas, S.; Hardcastle, C.; Green, D.; Ramsden, S.; Birney, E. J.; Black, G. C.; Sergouniotis, P.
Show abstract
PurposeThe PAX6 gene encodes a highly-conserved transcription factor involved in eye development. Heterozygous loss-of-function variants in PAX6 can cause a range of ophthalmic disorders including aniridia. A key molecular diagnostic challenge is that many PAX6 missense changes are presently classified as variants of uncertain significance. While computational tools can be used to assess the effect of genetic alterations, the accuracy of their predictions varies. Here, we evaluated and optimised the performance of computational prediction tools in relation to PAX6 missense variants. MethodsThrough inspection of publicly available resources (including HGMD, ClinVar, LOVD and gnomAD), we identified 241 PAX6 missense variants that were used for model training and evaluation. The performance of ten commonly-used computational tools was assessed and a threshold optimization approach was utilized to determine optimal cut-off values. Validation studies were subsequently undertaken using PAX6 variants from a local database. ResultsAlphaMissense, SIFT4G and REVEL emerged as the best-performing predictors; the optimized thresholds of these tools were 0.967, 0.025, and 0.772, respectively. Combining the prediction from these top-three tools resulted in lower performance compared to using AlphaMissense alone. ConclusionTailoring the use of computational tools by employing optimized thresholds specific to PAX6 can enhance algorithmic performance. Our findings have implications for PAX6 variant interpretation in clinical settings.
Khan, M.; Cornelis, S. S.; Pozo-Valero, M. d.; Whelan, L.; Runhart, E. H.; Mishra, K.; Bults, F.; AlSwaiti, Y.; AlTabishi, A.; Baere, E. D.; Banfi, S.; Banin, E.; Bauwens, M.; Ben-Yosef, T.; Boon, C. J. F.; Born, L. I. v. d.; Defoort, S.; Devos, A.; Dockery, A.; Dudakova, L.; Fakin, A.; Farrar, G. J.; Ferraz Sallum, J. M.; Fujinami, K.; Gilissen, C.; Glavac, D.; Gorin, M. B.; Greenberg, J.; Hayashi, T.; Hettinga, Y.; Hoischen, A.; Hoyng, C. B.; Hufendiek, K.; Jagle, H.; Kamakari, S.; Karali, M.; Kellner, U.; Klaver, C. C. W.; Kousal, B.; Lamey, T.; MacDonald, I. M.; Matynia, A.; McLaren, T.; M
Show abstract
Missing heritability in human diseases represents a major challenge. Although whole-genome sequencing enables the analysis of coding and non-coding sequences, substantial costs and data storage requirements hamper its large-scale use to (re)sequence genes in genetically unsolved cases. The ABCA4 gene implicated in Stargardt disease (STGD1) has been studied extensively for 22 years, but thousands of cases remained unsolved. Therefore, single molecule molecular inversion probes were designed that enabled an automated and cost-effective sequence analysis of the complete 128-kb ABCA4 gene. Analysis of 1,054 unsolved STGD and STGD-like probands resulted in bi-allelic variations in 448 probands. Twenty-seven different causal deep-intronic variants were identified in 117 alleles. Based on in vitro splice assays, the 13 novel causal deep-intronic variants were found to result in pseudo-exon (PE) insertions (n=10) or exon elongations (n=3). Intriguingly, intron 13 variants c.1938-621G>A and c.1938-514G>A resulted in dual PE insertions consisting of the same upstream, but different downstream PEs. The intron 44 variant c.6148-84A>T resulted in two PE insertions that were accompanied by flanking exon deletions. Structural variant analysis revealed 11 distinct deletions, two of which contained small inverted segments. Uniparental isodisomy of chromosome 1 was identified in one proband. Integrated complete gene sequencing combined with transcript analysis, identified pathogenic deep-intronic and structural variants in 26% of bi-allelic cases not solved previously by sequencing of coding regions. This strategy serves as a model study that can be applied to other inherited diseases in which only one or a few genes are involved in the majority of cases.
Yates, T. M.; Ansari, M.; Thompson, L.; Hunt, S. E.; Cibrian Uhalte, E.; Hobson, R. J.; Marsh, J. A.; Wright, C. F.; Firth, H. V.
Show abstract
Genetically determined disorders are highly heterogenous in clinical presentation and underlying molecular mechanism. The evidence underpinning these conditions in the peer-reviewed literature is variable and requires robust critical evaluation for diagnostic use. Here, we present a structured curation process for the Gene2Phenotype (G2P) project. This draws on multiple lines of clinical, bioinformatic and functional evidence. The process utilises and extends existing terminologies, allows for precise definition of the molecular basis of disease, and confidence levels to be attributed to a given gene-disease assertion. In-depth disease curation using this process will prove useful in applications including in diagnostics, research and the development of targeted therapeutics.
Hiramuki, Y.; Kure, Y.; Saito, Y.; Ogawa, M.; Ishikawa, K.; Mori-Yoshimura, M.; Oya, Y.; Takahashi, Y.; Kim, D.-S.; Arai, N.; Mori, C.; Matsumura, T.; Hamano, T.; Nakamura, K.; Ikezoe, K.; Hayashi, S.; Goto, Y.; Noguchi, S.; Nishino, I.
Show abstract
Facioscapulohumeral muscular dystrophy (FSHD) can be subdivided into two types: FSHD1, caused by contraction of the D4Z4 repeat on chromosome 4q35, and FSHD2, caused by mild contraction of the D4Z4 repeat plus aberrant hypomethylation mediated by genetic variants in SMCHD1, DNMT3B, or LRIF1. Genetic diagnosis of FSHD is challenging because of the complex procedures required. Here, we applied Nanopore CRISPR/Cas9-targeted resequencing for the diagnosis of FSHD by simultaneous detection of D4Z4 repeat length and methylation status at nucleotide level in genetically-confirmed and suspected patients. We found significant hypomethylation of contracted D4Z4 repeats in FSHD1 and strong correlation between methylation rate and patient phenotype. This finding can explain how repeat contraction contributes to disease pathogenesis by activating DUX4 expression.
Davydenko, K.; Skoblov, M.; Filatova, A.
Show abstract
BackgroundPathogenic variants in the dystrophin (DMD) gene lead to X-linked recessive Duchenne muscular dystrophy (DMD) and Becker muscular dystrophy (BMD). Nucleotide variants that affect splicing are a known cause of hereditary diseases. However, their representation in the public genomic variation databases is limited due to the low accuracy of their interpretation, especially if they are located within exons. The analysis of splicing variants in the DMD gene is essential both for understanding the underlying molecular mechanisms of the dystrophinopathies pathogenesis and selecting suitable therapies for patients. ResultsUsing deep in silico mutagenesis of the entire DMD gene sequence and subsequent SpliceAI splicing predictions, we identified 7,948 DMD single nucleotide variants that could potentially affect splicing, 863 of them were located in exons. Next, we analyzed over 1,300 disease-associated DMD SNVs previously reported in the literature (373 exonic and 956 intronic) and intersected them with SpliceAI predictions. We predicted that [~]95% of the intronic and [~]10% of the exonic reported variants could actually affect splicing. Interestingly, the majority (75%) of patient-derived intronic variants were located in the AG-GT terminal dinucleotides of the introns, while these positions accounted for only 13% of all intronic variants predicted in silico. Of the 97 potentially spliceogenic exonic variants previously reported in patients with dystrophinopathy, we selected 38 for experimental validation. For this, we developed and tested a minigene expression system encompassing 27 DMD exons. The results showed that 35 (19 missense, 9 synonymous, and 7 nonsense) of the 38 DMD exonic variants tested actually disrupted splicing. We compared the observed consequences of splicing changes between variants leading to severe Duchenne and milder Becker muscular dystrophy and showed a significant difference in their distribution. This finding provides extended insights into relations between molecular consequences of splicing variants and the clinical features. ConclusionsOur comprehensive bioinformatics analysis, combined with experimental validation, improves the interpretation of splicing variants in the DMD gene. The new insights into the molecular mechanisms of pathogenicity of exonic single nucleotide variants contribute to a better understanding of the clinical features observed in patients with Duchenne and Becker muscular dystrophy.
van der Sanden, B.; Neveling, K.; Shukor, S.; Gallagher, M. D.; Lee, J.; Burke, S. L.; Pennings, M.; van Beek, R.; Oorsprong, M.; Kater-Baats, E.; Kamping, E.; Tieleman, A.; Voermans, N.; Scheffer, I. E.; Gecz, J.; Corbett, M.; Vissers, L. E.; Pang, A. W.; Hastie, A.; Kamsteeg, E.-J.; Hoischen, A.
Show abstract
Short tandem repeats (STRs) are amongst the most abundant class of variations in human genomes and are meiotically and mitotically unstable which leads to expansions and contractions. STR expansions are frequently associated with genetic disorders, with the size of expansions often correlating with the severity and age of onset. Therefore, being able to accurately detect the total repeat expansion length and to identify potential somatic repeat instability is important. Current standard of care (SOC) diagnostic assays include laborious repeat-primed PCR-based tests as well as Southern blotting, which are unable to precisely determine long repeat expansions and/or require a separate set-up for each locus. Sequencing-based assays have proven their potential for the genome-wide detection of repeat expansions but have not yet replaced these diagnostic assays due to their inaccuracy to detect long repeat expansions (short-read sequencing) and their costs (long-read sequencing). Here, we tested whether optical genome mapping (OGM) can efficiently and accurately identify the STR length and assess the stability of known repeat expansions. We performed OGM for 85 samples with known clinically relevant repeat expansions in DMPK, CNBP and RFC1, causing myotonic dystrophy type 1 and 2 and cerebellar ataxia, neuropathy and vestibular areflexia syndrome (CANVAS), respectively. After performing OGM, we applied three different repeat expansion detection workflows, i.e. manual de novo assembly, local guided assembly (local-GA) and molecule distance script of which the latter two were developed as part of this study. The first two workflows estimated the repeat size for each of the two alleles, while the third workflow was used to detect potential somatic instability. The estimated repeat sizes were compared to the repeat sizes reported after the SOC and concordance between the results was determined. All except one known repeat expansions above the pathogenic repeat size threshold were detected by OGM, and allelic differences were distinguishable, either between wildtype and expanded alleles, or two expanded alleles for recessive cases. An apparent strength of OGM over current SOC methods was the more accurate length measurement, especially for very long repeat expansion alleles, with no upper size limit. In addition, OGM enabled the detection of somatic repeat instability, which was detected in 9/30 DMPK, 23/25 CNBP and 4/30 RFC1 samples, leveraging the analysis of intact, native DNA molecules. In conclusion, for tandem repeat expansions larger than [~]300 bp, OGM provides an efficient method to identify exact repeat lengths and somatic repeat instability with high confidence across multiple loci simultaneously, enabling the potential to provide a significantly improved and generic genome-wide assay for repeat expansion disorders.
Aydin, E.; Ergun, B.; Akgun-Dogan, O.; Alanay, Y.; Hatirnaz Ng, O.; Ozdemir, O.
Show abstract
The clinical interpretation of missense variants is critically important in diagnostics due to their potential to cause mild-to-severe effects on phenotype by altering protein structure. Evaluating these variants is essential because they can significantly impact disease outcomes and patient management. Many computational predictors, known as in silico pathogenicity predictors (ISPPs), have been developed to support the assessment of variant pathogenicity. Despite the abundance of these ISPPs, their predictions often lack accuracy and consistency, primarily due to limited data availability and the presence of erroneous data. This inconsistency can lead to false positive or negative results in pathogenicity evaluation, highlighting the need for standardization. The necessity for reliable evaluation methods has driven the development of numerous ISPPs, each attempting to address different aspects of variant interpretation. However, the sheer number of ISPPs and their varied performances make it challenging to achieve consensus in predictions. Therefore, a comprehensive statistical approach to evaluate and integrate these predictors is essential to improve accuracy. Here, we present a comprehensive statistical analysis comparing 52 available ISPPs, which aims to enhance the precision of variant classification. Our work introduces the Variant Analysis with Multiple Pathogenicity Predictors-score (VAMPP-score), a novel statistical framework designed for the assessment of missense variants. The VAMPP-score leverages the best gene-ISPP matches based on ISPP accuracies, providing a combinatorial weighted score that improves missense variant interpretation. We chose to develop a statistical framework rather than creating a new ISPP to capitalize on the strengths of existing predictors and to address their limitations through an integrative approach. This approach not only improves the evaluation of missense variants but also offers a flexible statistical framework designed to identify and utilize the best-performing ISPPs. By enhancing the accuracy of genetic diagnostics, particularly in the reanalysis of rare and undiagnosed cases, our framework aims to improve patient outcomes and advance the field of genetic research. Our study employed a comprehensive workflow (Figure 1) to enhance the accuracy of genomic variant interpretation with in-silico pathogenicity predictor (ISPP) evaluation. This workflow led to three pivotal results: O_FIG O_LINKSMALLFIG WIDTH=171 HEIGHT=200 SRC="FIGDIR/small/602867v1_fig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@fc41f4org.highwire.dtl.DTLVardef@14e1d8eorg.highwire.dtl.DTLVardef@1765c75org.highwire.dtl.DTLVardef@1b00719_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOFigure 1.C_FLOATNO Methodological summary. A database was built using dbNSFP (1, 2) (v4.7A) data, which includes rank scores from over fifty computational predictors for all possible single nucleotide variations in the human genome. This database was enhanced with canonical transcript information from Ensembl. The analysis focused on missense variants in ClinVar (3), categorized as Pathogenic (P), Benign (B), and Unknown (U). A novel statistical framework was used to assess the gene-based performance of each ISPP in distinguishing these variant groups through multiple and pairwise comparisons and proximity estimations. ISPPs were categorized (Category A-D) and assigned weights based on their performance for each gene, known as gene coefficients. These were integrated into the final assessment, the VAMPP-score. C_FIG[bullet] ISPPs were categorized on their prediction approaches. This classification not only streamlined the analytical process but also enhanced the interpretability of predictor outputs. [bullet]Leveraging this categorization, we conducted a robust statistical analysis to evaluate the prediction accuracy and performance of each ISPP. Our findings revealed a significant correlation between the prediction approaches of the ISPPs and their predictive successes, confirming the utility of our categorization approach. [bullet]These insights enabled us to develop a novel scoring system--the VAMPP-score--which integrates ISPPs according to their performances.
van den Berg, R. R.; Lauffer, M. C.; Laros, J. F. J.
Show abstract
Visualization of genes and genetic variants as well as transcript structure is essential within the human genetics community. Such illustrations represent a key tool in communicating genetic concepts and facilitating discussions on therapeutic interventions. There currently are no easily usable tools which allows the users to draw all features required for a comprehensive overview of a transcripts structure and the localisation of variants of interest. Here we introduce ExonViz, an online application that creates biologically accurate transcript figures, including features such as coding regions, genetic variants and exon reading frames. Transcript definitions are automatically retrieved from Ensembl and RefSeq. We illustrate the full functionality of ExonViz by generating a figure for all variants reported in ClinVar for CYLD. ExonViz is available online via the Dutch Center for RNA Therapeutics website and can be installed locally via PyPI.
Silverstein, S.; Orbach, R.; Syeda, S.; Foley, R.; Gorokhova, S.; Meilleur, K. G.; Leach, M. E.; Uapinyoying, P.; Chao, K. R.; Donkervoort, S.; Bönnemann, C. G.
Show abstract
Biallelic pathogenic variants in the gene encoding nebulin (NEB) are a known cause of congenital myopathy. We present two individuals with congenital myopathy and compound heterozygous variants (NM_001271208.2: c.2079C>A; p.(Cys693Ter) and c.21522+3A>G) in NEB. Transcriptomic sequencing on patient muscle revealed that the extended splice variant c.21522+3A>G causes exon 144 skipping. Nebulin isoforms containing exon 144 are known to be mutually exclusive with isoforms containing exon 143, and these isoforms are differentially expressed during development and in adult skeletal muscles. Patients MRIs were compared to the known pattern of relative abundance of these two isoforms in muscle. We propose that the pattern of muscle involvement in these patients better fits the distribution of exon 144-containing isoforms in muscle than with previously published MRI findings in NEB-related disease due to other variants. To our knowledge this is the first report hypothesizing disease pathogenesis through the alteration of isoform distributions in muscle.
Youmans, L.; Kamath, C.; Mansoorshahi, S.; Kurjee, M.; Laville, P.; Sprenger, A.; Frost, J.; Miller, R.; Northrup, H.; Au, K. S.
Show abstract
Myelomeningocele (MMC) is the most severe form of an open neural tube defect (NTD) that is compatible with life. The prevalence of MMC in the United States is 1 in 2,500 live births, with the two ethnicities that have the highest occurrence of MMC being Mexican American (MA) and Caucasian American (EA). Research to date has shown that MMC results from a cumulative effect of environmental and genetic factors. Therefore, determining the underlying molecular etiology would be a step toward developing strategies for prevention and treatment. We examined variants in 568 nervous system development genes implicated in MMC by whole exome sequencing of 254 MA and 257 EA subjects born with MMC. Mutational burden analysis was used to compare the deleterious variant load between MMC subjects and the reference population in the Genome Aggregation Exome Database (gnomADe). Higher mutational burdens were found in 18 genes, with PTK2 being the most significant (OR=3.49, p=5.3e-3) among genes known to be expressed in the human neural tube at the CS12/CS13 stages. Cell migration assay was performed using seven PTK2 (aka FAK1) deleterious variants in transfected Fak-/- mouse embryonic cells. The effect of Ptk2 knockdown on neural tube development was examined using Xenopus embryos. Cell migration assay results showed the seven MMC-associated PTK2 variants significantly affected migratory capacity compared to the wild-type PTK2. The Knockdown of Ptk2 significantly affected the normal neural tube closure of Xenopus embryos. Based on these findings, PTK2 variants identified from MMC patients may play a role in the multifactorial causation of MMC.
Yin, X.; Richardson, M. E.; Laner, A.; Shi, X.; Ognedal, E.; Vasta, V.; Hansen, T. v. O.; Pienda, M.; Ritter, D.; den Dunnen, J. T.; Hassanin, E.; Lyman Lin, W.; Borras, E.; Krahn, K.; Nordling, M.; Martins, A.; Mahmood, K.; Nadeau, E. A. W.; Beshay, V.; Tops, C.; Genuardi, M.; Pesaran, T.; Frayling, I. M.; Capella, G.; Latchford, A.; Tavtigian, S. V.; Maj, C.; Plon, S. E.; Greenblatt, M. S.; Macrae, F. A.; Spier, I.; Aretz, S.
Show abstract
BackgroundPathogenic constitutional APC variants underlie familial adenomatous polyposis, the most common hereditary gastrointestinal polyposis syndrome. To improve variant classification and resolve the interpretative challenges of variants of uncertain significance (VUS), APC-specific ACMG/AMP variant classification criteria were developed by the ClinGen-InSiGHT Hereditary Colorectal Cancer/Polyposis Variant Curation Expert Panel (VCEP). MethodsA streamlined algorithm using the APC-specific criteria was developed and applied to assess all APC variants in ClinVar and the InSiGHT international reference APC LOVD variant database. ResultsA total of 10,228 unique APC variants were analysed. Among the ClinVar and LOVD variants with an initial classification of (Likely) Benign or (Likely) Pathogenic, 94% and 96% remained in their original categories, respectively. In contrast, 41% ClinVar and 61% LOVD VUS were reclassified into clinically actionable classes, the vast majority as (Likely) Benign. The total number of VUS was reduced by 37%. In 21 out of 36 (58%) promising APC variants that remained VUS despite evidence for pathogenicity, a data mining-driven work-up allowed their reclassification as (Likely) Pathogenic. ConclusionsThe application of APC-specific criteria substantially reduced the number of VUS in ClinVar and LOVD. The study also demonstrated the feasibility of a systematic approach to variant classification in large datasets, which might serve as a generalisable model for other gene-/disease-specific variant interpretation initiatives. It also allowed for the prioritization of VUS that will benefit from in-depth evidence collection. This subset of APC variants was approved by the VCEP and made publicly available through ClinVar and LOVD for widespread clinical use.
Bohn, E.; Lau, T.; Wagih, O.; Masud, T.; Merico, D.
Show abstract
Variants in 5 and 3 untranslated regions (UTR) contribute to rare disease. While predictive algorithms to assist in classifying pathogenicity can potentially be highly valuable, the utility of these tools is often unclear, as it depends on carefully selected training and validation conditions. To address this, we developed a high-confidence set of pathogenic (P) and likely pathogenic (LP) variants and assessed deep learning (DL) models for predicting their molecular effect. 3 and 5 UTR variants documented as P or LP (P/LP) were obtained from ClinVar and refined by reviewing the annotated variant effect and reassessing evidence of pathogenicity following published guidelines. Prediction scores from sequence-based DL models were compared between three groups: P/LP variants acting though the mechanism for which the model was designed (model-matched), those operating through other mechanisms (model-mismatched), and putative benign variants. PhyloP was used to compare conservation scores between P/LP and putative benign variants. 295 3 and 188 5 UTR variants were obtained from ClinVar, of which 26 3 and 68 5 UTR variants were classified as P/LP. Predictions by DL models achieved statistically-significant differences when comparing model-matched P/LP variants to both putative benign variants and model-mismatched P/LP variants, as well as when comparing all P/LP variants to putative benign variants. PhyloP conservation scores were significantly higher among P/LP compared to putative benign variants for both the 3 and 5 UTR. In conclusion, we present a high-confidence set of P/LP 3 and 5 UTR variants spanning a range of mechanisms and supported by detailed pathogenicity and molecular mechanism evidence curation. Predictions from DL models further substantiate these classifications. These datasets will support further development and validation of DL algorithms designed to predict the functional impact of variants that may be implicated in rare disease.
Eisenhart, C.; Mewton, J.; Brickey, R.; Bayat, V.
Show abstract
In 2015, the American College of Medical Genetics and Genomics (ACMG) in collaboration with the Association of Molecular Pathologists (AMP) published guidelines for the interpretation and classification of germline genomic variants. The ACMG terminology guidelines outlined criteria for assigning one of five categories: benign, likely benign, uncertain significance, likely pathogenic and pathogenic. While the paper laid out 28 different classifiers and the justification for them, it did not provide specific algorithms for implementing these classifiers in an automated manner. Here we present the Bitscopic Interpreting ACMG Standards 2015 (BIAS-2015) software as a complete, open-source algorithm which categorizes variants according to the ACMG classification system. BIAS-2015 evaluates 18 of the 28 ACMG criteria to classify variants in an automated and consistent way while recording the rationale for each classifier to enable in-depth review. We used the genomic data from the ClinGen Evidence Repository (eRepo v1.0.29), one of two FDA-recognized human genetic variant databases, to evaluate the performance of the BIAS-2015 algorithm. All code for BIAS-2015 has been made available on GitHub.
Livesey, B. J.; Marsh, J. A.
Show abstract
Computational variant effect predictors (VEPs) and multiplexed assays of variant effect (MAVEs) are two key tools for assessing the functional consequences of genetic variants. While their outputs are often concordant, there are also many differences. Here, we analyse missense MAVE data from 37 different human proteins, comparing them to five state-of-the-art VEPs in order to quantify and explain their points of agreement and disagreement. We find that discordance is not random but reflects fundamental differences in how each method infers functional impact. VEPs, which rely heavily on sequence conservation and basic structural features, tend to overcall pathogenicity at buried and hydrophobic residues, while underestimating impact in disordered regions and on charged surface residues. MAVEs, by contrast, capture context-specific mechanisms more accurately, but can miss pathogenic variants when the assay fails to reflect disease biology, or be subject to high levels of experimental noise. By comparing both global patterns and specific clinically relevant variants, we show how protein features, assay design, and variant type shape predictive discordance. Our findings provide a framework for interpreting when and why VEPs and MAVEs diverge and point toward strategies for improving variant interpretation through integrated, mechanism-aware approaches.
Aissi, D.; Soukarieh, O.; Proust, C.; Jaspard-Vinassa, B.; Fautrad, P.; Ibrahim-Kosta, M.; Leal-Valentim, F.; Roux, M.; Bacq-Daian, D.; Olaso, R.; Deleuze, J.-F.; Morange, P.-E.; Tregouet, D.-A.
Show abstract
SummaryVariants in 5UTR regions that create upstream translation initiation AUG codons are a class of neglected non coding variations. When they associate with a premature stop codon and create upstream open reading frames (uORFs) whose translation competes with that of natural proteins, they can have strong impact on human diseases. We here describe MORFEE, a new bioinformatics tool that detects, annotates and predicts, from a standard VCF file, the creation of uORF by any 5UTR variants on uORF creation. MORFEE was applied to two genomic resources and identified candidate functional variants that could explain statistical association signals observed in the context of Genome Wide Association Studies or could be responsible for rare forms of diseases. In conclusion MORFEE is an easy-to-use tool complementary to existing ones that can help resolving genetic investigations that remained so far unfruitful. Availability and implementationMORFEE is written in R with code and package available at https://github.com/daissi/MORFEE. Contactdavid-alexandre.tregouet@inserm.fr; david-alexandre.tregouet@u-bordeaux.fr
Mondal, S.; Dutta, A. K.; Goswami, K.
Show abstract
While many Rare Inborn Errors of Metabolism are treatable conditions their optimal diagnosis and treatment is a challenge for nations with low resources. Moreover, the population prevalence of these conditions is largely unknown. The availability of large genomic datasets brings the opportunity to estimate population carrier frequency of autosomal recessive IEMs. This would help to generate diseases burden statistics for better allocation of resources. In the current work we estimated the gene specific combined minor allele frequency of pathogenic variants from the gnomAD dataset for 235 genes associated with IEM phenotypes in OMIM. As per our estimation almost one third of the Global population is carrier for a pathogenic variant responsible for rare autosomal recessive inborn error of metabolism with the highest carrier frequency in the Ashkenazi Jews. Globally per thousand live births approximately five children are born with an ARIEM. European Finnish have the highest burden of nine out of 10,000 live births. With 25 million live births per year India is expected to have at least 8,025 newborns with an ARIEM. Since many of these diseases are treatable early newborn screening holds the key to ensure optimal management of these children.
Eichstaedt, C. A.; Maldonado-Velez, G.; Machado, R. D.; Balachandar, S.; Coulet, F.; Day, K.; Dooijes, D.; Eyries, M.; Graef, S.; Macaya, D.; Shaukat, M.; Southgate, L.; Tenorio-Castano, J.; Chung, W. K.; Welch, C. L.; Aldred, M. A.
Show abstract
Purpose: Pulmonary arterial hypertension (PAH) is a rare disease that can be caused by pathogenic variants, most frequently in the bone morphogenetic protein receptor type 2 (BMPR2) gene. We formed a ClinGen variant curation expert panel to devise guidelines for the clinical interpretation of BMPR2 variants identified in PAH patients. Methods: The general ACMG/AMP variant classification criteria were refined for PAH and adapted to BMPR2 following ClinGen procedures. Subsequently, these specifications were tested independently by three members of the curation expert panel on 28 representative BMPR2 variants selected from ClinVar, and then presented and discussed in the plenum. Results: Application of the final BMPR2 variant specifications resolved 6 of 9 variants (66%) where multiple ClinVar classifications included a Variant of Uncertain Significance, with all six being reclassified as Benign or Likely Benign. Four splice site variants underwent clinically consequential reclassifications based on the presence or absence of supporting mRNA splicing data. Conclusion: The variant specifications provide an international framework and a useful tool for BMPR2 variant classification and can be applied to increase confidence and consistency in BMPR2 interpretation for diagnostic laboratories, clinical providers, and patients.
Ranjan, P.; Devi, C.; Verma, N.; Bansal, R.; Srivastava, V. K.; Das, P.
Show abstract
This study investigates the genetic underpinnings of congenital tooth agenesis (CTA) using a multi-omics approach, integrating whole exome sequencing (WES)and RNA expression analysis. WES was used to analyze the genetic basis of CTA in six affected individuals one with syndromic and five with non-syndromic CTA alongside three healthy and two internal controls. We identified both known and novel variants in candidate genes (EDA, WNT10A, PAX9, TSPEAR) and assessed the functional impacts of novel variants (WNT10A (A135S), and compound heterozygous TSPEAR (L219P, I419Lfs*150) using RT-PCR, while bioinformatics tools were applied to both known and novel variants. RT-PCR indicated disrupted EDA and WNT10A signaling in novel candidate genes WNT10A and TSPEAR. Computational analysis showed deleterious effects for six variants, with gene ontology, protein disorder, localization, and post-translational modifications suggesting significant functional changes. Molecular dynamics simulations predicted that these variants could impact protein stability and function. Additionally, WES analysis revealed 21 genes consistently present in all patients (MAF [≤]20%), including novel variants in OR4F21 (K310R, F44L) and LCORL (L1734P). Two variants, OR4F21 (K310R) and MRTFB (A135A), appeared in all cases. Furthermore, 391 genes were shared among three patients, 204 among four, and 98 among five. Integrating multi-omic data from the GEO database identified 18 upregulated and 15 downregulated genes, with variants linked to systemic conditions such as autism, Alzheimers, congenital heart disease (CHD), ALS (amyotrophic lateral sclerosis) and cancer. Our findings provide insights into CTAs molecular mechanisms, identifying potential biomarkers and therapeutic targets. Further validation could improve diagnosis and treatment strategies for CTA. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=155 SRC="FIGDIR/small/24317461v1_ufig1.gif" ALT="Figure 1"> View larger version (93K): org.highwire.dtl.DTLVardef@718fe2org.highwire.dtl.DTLVardef@19d70ecorg.highwire.dtl.DTLVardef@1609ceforg.highwire.dtl.DTLVardef@1df3dc7_HPS_FORMAT_FIGEXP M_FIG C_FIG
Ahmed P, H.; Singh, P.; Thakur, R.; Kumari, A.; Krishnan, H.; Philip, R. G.; Vasudevan, A.; PADINJAT, R.
Show abstract
Lowe syndrome is an X-linked recessive monogenic disorder resulting from mutations in the OCRL gene that encodes a phosphatidylinositol 4,5 bisphosphate 5-phosphatase. The disease affects three organs-the kidney, brain and eye and clinically manifests as proximal renal tubule dysfunction, neurodevelopmental delay and congenital cataract. Although Lowe syndrome is a monogenic disorder, there is considerable heterogeneity in clinical presentation; some individuals show primarily renal symptoms with minimal neurodevelopmental impact whereas others show neurodevelopmental defect with minimal renal symptoms. However, the molecular and cellular mechanisms underlying this clinical heterogeneity remain unknown. Here we analyze a Lowe syndrome family in whom affected members show clinical heterogeneity with respect to the neurodevelopmental phenotype despite carrying an identical mutation in the OCRL gene. Genome sequencing and variant analysis in this family identified a large number of damaging variants in each patient. Using novel analytical pipelines and segregation analysis we prioritize variants uniquely present in the patient with the severe neurodevelopmental phenotype compared to those with milder clinical features. The identity of genes carrying such variants underscore the role of additional gene products enriched in the brain or highly expressed during brain development that may be determinants of the neurodevelopmental phenotype in Lowe syndrome. We also identify a heterozygous variant in CEP290, previously implicated in ciliopathies that underscores the potential role of OCRL in regulating ciliary function that may impact brain development. More generally, our findings demonstrate analytic approaches to identify high-confidence genetic variants that could underpin the phenotypic heterogeneity observed in monogenic disorders.